Ir arriba
Información del artículo en conferencia

SemTab: A Hybrid Framework for Semantic Feature Generation on Tabular Data

O. Chen, K. Chou, R. Nagpal, R. Palacios, A. Gupta

Undergraduate Research Technology Conference - MIT URTC 2025, Cambridge (Estados Unidos de América). 10-12 octubre 2025


Resumen:

Machine learning models on tabular datasets often struggle to understand the context between features, which can limit their accuracy. We propose SemTab, a hybrid framework for generating semantic features that utilizes an open-source Large Language Model (LLM). We evaluated our framework using three benchmark datasets: Adult Income, German Credit, and Bank Marketing. We compared its performance against several off-the-shelf LLMs. The results show that SemTab achieved the highest accuracy across all the classification tasks. For instance, on the Bank Marketing dataset, SemTab achieved an accuracy of 8 0%, which is approximately 2 0% improvement over the baseline models. This work highlights that a hybrid architecture is a practical approach for applying language models to structured tabular data, yielding accurate and interpretable results for various downstream tasks.


Palabras clave: Tabular Data, Semantic Feature Generation, LLMs, Model Interpretability


DOI: DOI icon https://doi.org/10.1109/URTC68753.2025.11533131

Publicado en: 2025 IEEE MIT Undergraduate Research Technology Conference (URTC), pp: 1-5, ISBN: 979-8-3315-5938-0

Fecha de publicación: 26-may-2026


Cita:
O. Chen, K. Chou, R. Nagpal, R. Palacios, A. Gupta, "SemTab: A Hybrid Framework for Semantic Feature Generation on Tabular Data", presentado en Undergraduate Research Technology Conference - MIT URTC 2025, Cambridge, Estados Unidos de América, 10-12 octubre 2025. En: 2025 IEEE MIT Undergraduate Research Technology Conference (URTC), pp. 1-5, doi: 10.1109/URTC68753.2025.11533131

    Grupos de investigación:
  • Instituto de Investigación Tecnológica (IIT)

IIT-25-413C

pdf Solicitar el artículo completo a los autores